Back

Genome Biology

Springer Science and Business Media LLC

All preprints, ranked by how well they match Genome Biology's content profile, based on 637 papers previously published here. The average preprint has a 0.47% match score for this journal, so anything above that is already an above-average fit. Older preprints may already have been published elsewhere.

1
glmmDMR reveals replicate-level methylation variance as a major determinant of false-positive DMR detection

Daito, Y.; Uechi, M.; Kinoshita, T.; Tonosaki, K.

2026-07-03 genomics 10.64898/2026.06.29.734667 medRxiv
Top 0.1%
61.1%
Show abstract

Background: Accurate identification of differentially methylated regions (DMRs) is fundamental to epigenomic research but remains challenging due to biological variability among replicates, heterogeneous effect sizes, and the tendency of adjacent cytosines to share similar methylation states. Many existing methods aggregate methylation measurements before statistical testing or do not explicitly account for replicate-level variability, contributing to elevated false-positive rates. Results: We developed glmmDMR, a DMR detection framework that combines generalized linear mixed models with a seed-based strategy for reconstructing DMRs from locally high-confidence signals while explicitly modeling replicate-level variability. Using simulated datasets with known ground-truth DMRs, we demonstrate that false-positive detections are more strongly associated with methylation variance among biological replicates than with the magnitude of methylation differences between groups. glmmDMR achieved higher precision than existing approaches while maintaining competitive recall, particularly for subtle methylation differences. Site-level modeling with beta regression provided the strongest overall performance, and seed-based region construction reduced artificial DMR fragmentation, improving recovery of true DMR boundaries and producing more contiguous, biologically interpretable DMRs. Applied to Arabidopsis thaliana ddm1 methylomes and a rice DEMETER-LIKE DNA demethylase mutant (Osdml3a-1), glmmDMR identified biologically meaningful DMRs, revealing widespread TE-associated hypomethylation and subtle TE-family-specific hypermethylation. Conclusions: Replicate-level methylation variance is an important determinant of DMR detection performance, and explicitly modeling this variance improves discrimination of biologically meaningful methylation changes from high-variance signals. By combining variance-aware statistical modeling with seed-based region construction, glmmDMR provides a robust framework for identifying contiguous, biologically interpretable DMRs across diverse methylome datasets.

2
Benchmarking alternative polyadenylation detection in single-cell and spatial transcriptomes

Li, S.; Wang, Z.; Hu, Y.; Ni, Q.; Feng, C.; Hu, Y.; Zhang, S.; Chen, M.

2024-10-17 bioinformatics 10.1101/2024.10.15.618405 medRxiv
Top 0.1%
61.0%
Show abstract

Background3-tag-based sequencing methods have become the predominant approach for single-cell and spatial transcriptomics, with some protocols proven effective in detecting alternative polyadenylation (APA). While numerous computational tools have been developed for APA detection from these sequencing data, the absence of comprehensive benchmarks and the diversity of sequencing protocols and tools make it challenging to select appropriate methods for APA analysis in these contexts. ResultsWe systematically compared seven 3-tag-based sequencing protocols and identified key peak features affecting APA detection performance. We developed a simulation pipeline that generates realistic datasets preserving protocol-specific characteristics. Using simulated and real data, we comprehensively assessed six computational tools for their ability to identify polyA sites, quantify polyA site expression, detect differentially expressed (DE) APA genes, filter sequencing artifacts, and their computational efficiency. We also investigated factors influencing APA detection. Our evaluation revealed that SCAPE and scAPAtrap generally outperformed other tools across various performance metrics and protocols. ConclusionOur systematic evaluation provides guidance for tool selection, experiment design, and future tool development in APA analysis for singlecell and spatial transcriptomics, paving the way for investigating APA in these contexts.

3
Experimental and Computational Methods for Allelic Imbalance Analysis from Single-Nucleus RNA-seq Data

Simmons, S. K.; Adiconis, X.; Haywood, N.; Parker, J.; Lin, Z.; Liao, Z.; Tuncali, I.; Al'Khafaji, A.; Shin, A.; Jagadeesh, K.; Gosik, K.; Gatzen, M.; Smith, J. T.; El Kodsi, D. N.; Kuras, Y.; Baecher-Allan, C.; Serrano, G. E.; Beach, T. G.; Garimella, K.; Rozenblatt-Rosen, O.; Regev, A.; Dong, X.; Scherzer, C.; Levin, J. Z.

2024-08-16 genomics 10.1101/2024.08.13.607784 medRxiv
Top 0.1%
58.4%
Show abstract

Single-cell RNA-seq (scRNA-seq) is emerging as a powerful tool for understanding gene function across diverse cells. Recently, this has included the use of allele-specific expression (ASE) analysis to better understand how variation in the human genome affects RNA expression at the single-cell level. We reasoned that because intronic reads are more prevalent in single-nucleus RNA-Seq (snRNA-Seq), and introns are under lower purifying selection and thus enriched for genetic variants, that snRNA-seq should facilitate single-cell analysis of ASE. Here we demonstrate how experimental and computational choices can improve the results of allelic imbalance analysis. We explore how experimental choices, such as RNA source, read length, sequencing depth, genotyping, etc., impact the power of ASE-based methods. We developed a new suite of computational tools to process and analyze scRNA-seq and snRNA-seq for ASE. As hypothesized, we extracted more ASE information from reads in intronic regions than those in exonic regions and show how read length can be set to increase power. Additionally, hybrid selection improved our power to detect allelic imbalance in genes of interest. We also explored methods to recover allele-specific isoform expression levels from both long- and short-read snRNA-seq. To further investigate ASE in the context of human disease, we applied our methods to a Parkinsons disease cohort of 94 individuals and show that ASE analysis had more power than eQTL analysis to identify significant SNP/gene pairs in our direct comparison of the two methods. Overall, we provide an end-to-end experimental and computational approach for future studies.

4
A Systematic Benchmark of High-Accuracy PacBio Long-Read RNA Sequencing for Transcript-Level Quantification

Wissel, D.; Mehlferber, M. M.; Nguyen, K. M.; Pavelko, V.; Tseng, E.; Robinson, M. D.; Sheynkman, G. M.

2025-06-30 bioinformatics 10.1101/2025.05.30.656561 medRxiv
Top 0.1%
55.7%
Show abstract

PacBio long-read RNA sequencing resolves transcripts with greater clarity than short-read technologies, yet its quantitative performance remains under-evaluated at scale. Here, we benchmark the high-throughput PacBio Kinnex platform against Illumina short-read RNA-seq using matched, deeply sequenced datasets across a time course of endothelial cell differentiation. Compared to Illumina, Kin-nex achieved comparable gene-level quantification and more accurate transcript discovery and transcript quantification. While Illumina detected more transcripts overall, many reflected potentially unstable or ambiguous estimates in complex genes. Kinnex largely avoids these issues, producing more reliable differential transcript expression (DTE) calls, despite a mild bias against short transcripts (shorter than 1.25 kb). When correcting Illumina for inferential variability, Kinnex and Illumina quantifications were highly concordant, demonstrating equivalent performance. We also benchmarked long-read tools, nominating Oarfish as the most efficient for our Kinnex data. Together, our results establish Kinnex as a reliable platform for full-length transcript quantification.

5
Differential Expression Analysis for Longitudinal Single-Cell RNA-Sequencing Studies Using REBEL

Wynn, E. A.; Mould, K. J.; Vestal, B. E.; Moore, C. M.

2026-05-11 genomics 10.64898/2026.05.06.723139 medRxiv
Top 0.1%
55.7%
Show abstract

Longitudinal scRNA-seq experiments offer a powerful approach for dissecting temporal gene expression dynamics in individual cell types. However, few methods have been developed specifically to address the unique statistical challenges of repeated measures in scRNA-seq data. Here, we introduce a novel method, REBEL (Repeated measures Empirical Bayes differential Expression analysis using Linear mixed models), for analyzing cell type-specific differential expression in repeated measures scRNA-seq experiments. Using simulation studies, we demonstrate that, relative to conventional repeated measures analysis methods and other scRNA-seq approaches, REBEL controls the false discovery rate and exhibits competitive power across a range of simulation scenarios. We further validate REBEL by analyzing a longitudinal scRNA-seq dataset from patients with B-cell lymphoma receiving chimeric antigen receptor (CAR)-T cell therapy. REBEL is implemented as an R package, available at https://github.com/ewynn610/REBEL.

6
BacSC: A general workflow for bacterial single-cell RNA sequencing data analysis

Ostner, J.; Kirk, T.; Olayo-Alarcon, R.; Thöming, J. G.; Rosenthal, A. Z.; Häussler, S.; Müller, C. L.

2024-06-27 bioinformatics 10.1101/2024.06.22.600071 medRxiv
Top 0.1%
54.4%
Show abstract

Bacterial single-cell RNA sequencing has the potential to elucidate within-population heterogeneity of prokaryotes, as well as their interaction with host systems. Despite conceptual similarities, the statistical properties of bacterial single-cell datasets are highly dependent on the protocol, making proper processing essential to tap their full potential. We present BacSC, a fully data-driven computational pipeline that processes bacterial single-cell data without requiring manual intervention. BacSC performs data-adaptive quality control and variance stabilization, selects suitable parameters for dimension reduction, neighborhood embedding, and clustering, and provides false discovery rate control in differential gene expression testing. We validated BacSC on a broad selection of bacterial single-cell datasets spanning multiple protocols and species. Here, BacSC detected subpopulations in Klebsiella pneumoniae, found matching structures of Pseudomonas aeruginosa under regular and low-iron conditions, and better represented subpopulation dynamics of Bacillus subtilis. BacSC thus simplifies statistical processing of bacterial single-cell data and reduces the danger of incorrect processing.

7
Atlas of nascent RNA transcripts reveals enhancer to gene linkages

Sigauke, R. F.; Sanford, L.; Maas, Z. L.; Jones, T.; Stanley, J. T.; Townsend, H. A.; Allen, M. A.; Dowell, R. D.

2023-12-08 genomics 10.1101/2023.12.07.570626 medRxiv
Top 0.1%
54.4%
Show abstract

Gene transcription is controlled and modulated by regulatory regions, including enhancers and promoters. These regions are abundant in unstable, non-coding bidirectional transcription. Using nascent RNA transcription data across hundreds of human samples, we identified over 800,000 regions containing bidirectional transcription. We then identify highly correlated transcription between bidirectional and gene regions. The identified correlated pairs, a bidirectional region and a gene, are enriched for disease associated SNPs and often supported by independent 3D data. We present these resources as an SQL database which serves as a resource for future studies into gene regulation, enhancer associated RNAs, and transcription factors.

8
Multimodal single-cell chromatin analysis with Signac

Stuart, T.; Srivastava, A.; Lareau, C.; Satija, R.

2020-11-10 genomics 10.1101/2020.11.09.373613 medRxiv
Top 0.1%
54.2%
Show abstract

The recent development of experimental methods for measuring chromatin state at single-cell resolution has created a need for computational tools capable of analyzing these datasets. Here we developed Signac, a framework for the analysis of single-cell chromatin data, as an extension of the Seurat R toolkit for single-cell multimodal analysis. Signac enables an end-to-end analysis of single-cell chromatin data, including peak calling, quantification, quality control, dimension reduction, clustering, integration with single-cell gene expression datasets, DNA motif analysis, and interactive visualization. Furthermore, Signac facilitates the analysis of multimodal single-cell chromatin data, including datasets that co-assay DNA accessibility with gene expression, protein abundance, and mitochondrial genotype. We demonstrate scaling of the Signac framework to datasets containing over 700,000 cells. AvailabilityInstallation instructions, documentation, and tutorials are available at: https://satijalab.org/signac/

9
Single-cell RNA-seq differential expression tests within a sample should use pseudo-bulk data of pseudo-replicates

Hafemeister, C.; Halbritter, F.

2023-04-12 bioinformatics 10.1101/2023.03.28.534443 medRxiv
Top 0.1%
53.7%
Show abstract

Single-cell RNA sequencing (scRNA-seq) has become a standard approach to investigate molecular differences between cell states. Comparisons of bioinformatics methods for the count matrix transformation (normalization) and differential expression (DE) analysis of these data have already highlighted recommendations for effective between-sample comparisons and visualization. Here, we examine two remaining open questions: (i) What are the best combinations of data transformations and statistical test methods, and (ii) how do pseudo-bulk approaches perform in single-sample designs? We evaluated the performance of 343 DE pipelines (combinations of eight types of count matrix transformations and ten statistical tests) on simulated and real-world data, in terms of precision, sensitivity, and false discovery rate. We confirm superior performance of pseudo-bulk approaches without prior transformation. For within-sample comparisons, we advise the use of three pseudo-replicates, and provide a simple R package DElegate to facilitate application of this approach.

10
Igniting full-length isoform analysis in single-cell and spatial RNA-seq data with FLAMESv2

Wang, C.; Prawer, Y. D. J.; Voogd, O.; Schuster, J.; Pasquali, C.; De Paoli-Iseppi, R.; Li, A.; Hallab, J.; Tian, L.; Peng, H.; David, M.; Du, M. R. M.; Velasco, S.; Garone, M. G.; Dong, X.; Zeglinski, K.; Pavan, C.; Law, K. C. L.; Abu-Bonsrah, K. D.; Hunt, C. P. J.; Parish, C.; Gouil, Q.; Thijssen, R.; Davidson, N. M.; Ritchie, M. E.; Clark, M. B.; You, Y.

2026-03-12 bioinformatics 10.1101/2025.10.19.683327 medRxiv
Top 0.1%
53.5%
Show abstract

Long-read single-cell RNA-sequencing enables the profiling of RNA isoform expression and alternative splicing at single cell resolution. However, diverse single-cell technologies and sparse isoform data demand flexible and accurate analysis tools. We introduce FLAMESv2, a highly modular and protocol-agnostic R/Bioconductor package for long-read single-cell RNA-seq data analysis. FLAMESv2 supports a wide range of single-cell and spatial protocols, is highly configurable, scales to allow multi-sample analysis and provides versatile visualisation and analysis outputs. We demonstrate its compatibility with both droplet-based and combinatorial barcoding single-cell methods, as well as spatial transcriptomics workflows. Benchmarking confirms FLAMESv2 achieves field-leading performance across key analysis tasks. Applying FLAMESv2 to in vitro differentiation of stem cells into neurons, we identify cell-types, differentiation trajectories, expression of annotated and novel isoforms and isoform expression diversity and heterogeneity within individual cells. FLAMESv2 provides a comprehensive, flexible approach to analysing long-read single-cell RNA-sequencing, unlocking this powerful methodology for RNA isoform characterisation.

11
Transgressive gene expression and methylation remodeling in an intraspecific hexaploid wheat hybrid

Ardaman, A.; Forgiarini, C.; Arunkumar, R.

2026-07-09 plant biology 10.64898/2026.06.29.735383 medRxiv
Top 0.1%
52.4%
Show abstract

Intraspecific hybridization in allopolyploid plant genomes has the potential to induce non-additive changes in gene expression and DNA cytosine methylation, partly through interactions among divergent parental subgenomes. However, the extent to which intraspecific hybridization reshapes gene expression, coordinates homoeolog regulation, and remodels methylation in higher-order polyploids remains poorly quantified. To address this, we sequenced seedling leaf transcriptomes and methylomes from two parental cultivars of hexaploid bread wheat (Triticum aestivum L.) and their hybrids. More than 40% of genes were differentially expressed between hybrids and parents, although many were not differentially expressed between the parents themselves, consistent with complex trans-regulatory effects in the hybrid genome. This effect was more pronounced for homoeologs whose relative expression differed between the parents. These expression shifts often occurred simultaneously across all three homoeologs within triads, reducing homoeolog expression bias (HEB) in the hybrids. CG methylation levels were similar between the parents and hybrids in regions of low genetic divergence and in transposable element (TE)-rich regions, whereas CG sites in gene-rich regions showed more additive inheritance (hybrids intermediate between parents), particularly when parental haplotypes were themselves divergent. TE and gene body methylation (gbM) was strongly conserved in parents and hybrids. gbM was associated with more balanced homoeolog expression and fewer non-additive expression changes. CHH methylation showed overdominance, whereas non-conserved CHG methylation was enriched in TE-rich regions, suggesting that non-CG remodeling may reflect parental differences in TE and small-RNA content. Our results show that intraspecific hybridization within a hexaploid species can generate non-additive changes in gene expression and DNA methylation in seedling leaf tissue, while the presence of homoeologous genes, parental HEB, parental genetic and methylation divergence, and genomic location have varying levels of influence on expression or methylation remodeling.

12
Deepurify: a multi-modal deep language model to remove contamination from metagenome-assembled genomes

ZOU, B.; Wang, J.; Ding, Y.; Zhang, Z.; Yufen, H.; Fang, X.; Cheung, K. C.; See, S.; Zhang, L.

2023-09-29 genomics 10.1101/2023.09.27.559668 medRxiv
Top 0.1%
51.8%
Show abstract

Metagenome-assembled genomes (MAGs) offer valuable insights into the exploration of microbial dark matter using metagenomic sequencing data. However, there is a growing concern that contamination in MAGs may significantly impact the downstream analysis results. Existing MAG decontamination methods heavily rely on marker genes but do not fully leverage genomic sequences. To address the limitations, we have introduced a novel decontamination approach named Deepurify, which utilizes a multi-modal deep language model employing contrastive learning to learn taxonomic similarities of genomic sequences. Deepurify utilizes inferred taxonomic lineages to guide the allocation of contigs into a MAG-separated tree and employs a tree traversal strategy for maximizing the total number of medium- and high-quality MAGs. Extensive experiments were conducted on two simulated datasets, CAMI I, and human gut metagenomic sequencing data. These results demonstrate that Deepurify significantly outperforms other decontamination methods.

13
Built on sand: the shaky foundations of simulating single-cell RNA sequencing data

Crowell, H. L.; Leonardo, S. X. M.; Soneson, C.; Robinson, M. D.

2022-01-26 bioinformatics 10.1101/2021.11.15.468676 medRxiv
Top 0.1%
51.8%
Show abstract

With the emergence of hundreds of single-cell RNA-sequencing (scRNA-seq) datasets, the number of computational tools to analyse aspects of the generated data has grown rapidly. As a result, there is a recurring need to demonstrate whether newly developed methods are truly performant - on their own as well as in comparison to existing tools. Benchmark studies aim to consolidate the space of available methods for a given task, and often use simulated data that provide a ground truth for evaluations. Thus, demanding a high quality standard for synthetically generated data is critical to make simulation study results credible and transferable to real data. Here, we evaluated methods for synthetic scRNA-seq data generation in their ability to mimic experimental data. Besides comparing gene- and cell-level quality control summaries in both one- and two-dimensional settings, we further quantified these at the batch- and cluster-level. Secondly, we investigate the effect of simulators on clustering and batch correction method comparisons, and, thirdly, which and to what extent quality control summaries can capture reference-simulation similarity. Our results suggest that most simulators are unable to accommodate complex designs without introducing artificial effects; they yield over-optimistic performance of integration, and potentially unreliable ranking of clustering methods; and, it is generally unknown which summaries are important to ensure effective simulation-based method comparisons.

14
End-repair causes methylation underestimation in cell-free DNA sequencing libraries

Groth, T. E.; Mishin, A. A.; Rao, V.; Tibet, R.; Troll, C. J.

2025-12-17 genomics 10.64898/2025.12.15.694439 medRxiv
Top 0.1%
51.8%
Show abstract

Cell-free DNA methylation sequencing provides insight into tissue of origin and chromatin structure. In some workflows, generating libraries includes end-repair. Using matched single-stranded and double-stranded libraries prepared from the same cfDNA extracts, we show that end-repair in double-stranded DNA libraries reduces globally inferred CpG methylation leading to decreased tissue of origin accuracy. Trimming read termini partially mitigates this bias but decreases coverage and removes fragmentomic information compared to single-stranded DNA libraries, which forego end-repair.

15
A Genomic Language Model for Chimera Artifact Detection in Nanopore Direct RNA Sequencing

Li, Y.; Wang, T.-Y.; Guo, Q.; Ren, Y.; Lu, X.; Cao, Q.; Yang, R.

2024-10-25 genomics 10.1101/2024.10.23.619929 medRxiv
Top 0.1%
51.6%
Show abstract

Chimera artifacts in nanopore direct RNA sequencing (dRNA-seq) can significantly distort transcriptome analyses, yet their detection and removal remain challenging due to limitations in existing basecalling models. We present Deep-Chopper, a genomic language model that precisely identifies and removes adapter sequences from base-called dRNA-seq long reads at single-base resolution, operating independently of raw signal or alignment information to effectively eliminate chimeric read artifacts. By removing these artifacts, DeepChopper substantially improves the accuracy of critical downstream analyses, such as transcript annotation and gene fusion detection, thereby enhancing the reliability and utility of nanopore dRNA-seq for transcriptomics research.

16
scMetaIntegrator: a meta-analysis approach to paired single-cell differential expression analysis

Ratnasiri, K.; Mach, S. N.; Blish, C. A.; Khatri, P.

2025-06-08 bioinformatics 10.1101/2025.06.04.657898 medRxiv
Top 0.1%
51.5%
Show abstract

Traditional differential gene expression methods are limited for analysis of single cell RNA-sequencing (scRNA-seq) studies that use paired repeated measures and matched cohort designs. Many existing approaches consider cells as independent samples, leading to high false positive rates while ignoring inherent sampling structures. Although pseudobulk methods address this, they ignore intra-sample expression variability and have higher false negatives rates. We propose a novel meta-analysis approach that accounts for biological replicates and cell variability in paired scRNA-seq data. Using both real and synthetic datasets, we show that our method, single-cell MetaIntegrator (https://github.com/Khatri-Lab/scMetaIntegrator), provides robust effect size estimates and reproducible p-values.

17
Protein k-mers enable assembly-free microbial metapangenomics

Reiter, T. E.; Pierce-Ward, N. T.; Irber, L. C.; Botvinnik, O.; Brown, C. T.

2022-06-27 bioinformatics 10.1101/2022.06.27.497795 medRxiv
Top 0.1%
51.3%
Show abstract

An estimated 2 billion species of microbes exist on Earth with orders of magnitude more strains. Microbial pangenomes are created by aggregating all genomes of a single clade and reflect the metabolic diversity of groups of organisms. As de novo metagenome analysis techniques have matured and reference genome databases have expanded, metapangenome analysis has risen in popularity as a tool to organize the functional potential of organisms in relation to the environment from which those organisms were sampled. However, the reliance on assembly and binning or on reference databases often leaves substantial portions of metagenomes unanalyzed, thereby underestimating the functional potential of a community. To address this challenge, we present a method for metapangenomics that relies on amino acid k-mers (kaa-mers) and metagenome assembly graph queries. To enable this method, we first show that kaa-mers estimate pangenome characteristics and that open reading frames can be accurately predicted from short shotgun sequencing reads using the previously developed tool orpheum. These techniques enable pangenomics to be performed directly on short sequencing reads. To enable metapangenome analysis, we combine these approaches with compact de Bruijn assembly graph queries to directly generate sets of sequencing reads for a specific species from a metagenome. When applied to stool metagenomes from an individual receiving antibiotics over time, we show that these approaches identify strain fluctuations that coincide with antibiotic exposure.

18
Real-paired single-cell/bulk RNA-seq benchmark and a practical protocol for accurate cell-type deconvolution in human BAL samples

Hu, Y.; Liu, Z.; Tsao, D.; Leung, J. M.; V. Gerayeli, F.; Li, X.; Shao, X.; Sin, D.; Zhang, X.

2026-01-14 bioinformatics 10.64898/2026.01.14.699304 medRxiv
Top 0.1%
50.7%
Show abstract

BackgroundPseudo-bulk RNA-seq, generated by aggregating single-cell profiles, is widely used for benchmarking deconvolution methods because it globally approximates bulk transcriptomes and provides known cell-type proportions as ground truth. However, pseudo-bulk inherits single-cell specific measurement properties, and the extent to which these differ from real bulk RNA-seq remains difficult to quantify in the absence of paired data. In practice, deconvolution studies also commonly rely on large external single-cell references, yet the value of small, protocol-matched in-study references has not been systematically evaluated under real bulk conditions. These gaps motivate a paired benchmark that jointly examines pseudo-bulk fidelity, reference design, and their consequences for deconvolution accuracy. ResultsWe establish a fully paired benchmarking framework using split-sample, donor-matched bulk and single-cell RNA-seq (scRNA-seq) from human bronchoalveolar lavage (BAL). Embedded within a reverse five-fold cross-validation design and evaluated on real bulk RNA-seq data, this framework benchmarks 15 deconvolution algorithms published in 2013-2025 across 3 cell-type resolutions and 5 single-cell references, including three published BAL datasets, a harmonized lung BAL atlas, and an in-study reference derived from paired aliquots. We show that real bulk and matched pseudo-bulk profiles exhibit systematic gene-level differences, identifying 557 reproducibly discordant genes (|log2 FC| > 1, FDR< 0.05 by LIMMA), including cell-type informative features. These discrepancies reflect technology-specific effects and violate the linear mixing assumption underlying most deconvolution methods. We demonstrate that a protocol-matched in-study reference constructed from only six donors consistently outperforms substantially larger external references, with the advantage that increases at finer cell-type resolution. Moreover, paired samples enable the identification and selective removal of discordant genes, which further improves deconvolution accuracy for many algorithms, particularly in high-resolution settings. These findings are robust across methods, references, and evaluation criteria and extend beyond compositional accuracy to improving recovery of disease-associated cell-type differences in clinical applications. ConclusionsOur study provides the first fully paired benchmark of transcriptomic deconvolution on real bulk RNA-seq data of human BAL samples and demonstrates that reference design and data compatibility are as influential as algorithm choice. Beyond benchmarking, we introduce a practical and cost-effective protocol for deconvolution studies: generate single-cell data for a minimal subset of bulk samples (pilot pairing), use these data to construct an in-study reference and identify discordant genes, and apply the resulting insights to the full cohort. This strategy requires limited additional experimental effort yet yields substantial gains in accuracy and stability, offering actionable guidance for future bulk RNA-seq deconvolution studies across tissues and platforms. One-sentence summarySystematic benchmarking reveals pseudo-bulk biases and provides practical fixes for accurate RNA deconvolution.

19
DIANA: Deep Learning Identification and Assessment of Ancient DNA

Duitama Gonzalez, C.; Lopopolo, M.; Nishimura, L.; Faure, R.; Duchene, S.

2026-04-10 bioinformatics 10.64898/2026.04.09.717429 medRxiv
Top 0.1%
50.4%
Show abstract

The field of ancient metagenomics provides insights into past microbiomes, but with a growing dataset size, methods that rely on reference databases have limited scope. Here, we introduce DIANA, a multi-task neural network that predicts key metadata categories from unitig abundances. Trained on 2,597 run accessions (1.72 Tbp of assembled unitig sequences), DIANA accurately identifies sample host (94.6%), community type (90.0%), and material (88.9%) on held-out test data and demonstrates robust generalisation on an independent validation set. A key innovation is DIANAs ability to perform semantic generalisation, correctly classifying samples with labels unseen during training -- such as novel subspecies -- to their appropriate parent categories. By leveraging both known and uncharacterized genomic sequences, DIANA provides a rapid, data-driven system for metadata validation and quality control, accelerating discovery in ancient metagenomics research.

20
Pan-cell type continuous chromatin state annotation of all IHEC epigenomes

Daneshpajouh, H.; Moghul, I.; Wiese, K. C.; Libbrecht, M. W.

2025-02-08 genomics 10.1101/2025.02.06.636950 medRxiv
Top 0.1%
48.6%
Show abstract

The International Human Epigenome Consortium has generated thousands of epigenomic datasets that mea-sure various biochemical activities in the genome, including transcription factor binding, histone modification, and DNA accessibility. Currently, the predominant methods for integrating these datasets to annotate regu-latory elements are segmentation and genome annotation (SAGA) algorithms. The majority of annotations by these methods are cell type-specific. However, as the number of profiled cell types has grown into the thousands, using thousands of cell type-specific chromatin state annotations proves undesirable for many applications. Here, we present a pan-cell type annotation that summarizes all IHEC epigenomes using the recently-developed method, epigenome-ssm.